Papers by Janet B. Pierrehumbert

6 papers
Stories that (are) Move(d by) Markets: A Causal Exploration of Market Shocks and Semantic Shifts across Different Partisan Groups (2025.findings-acl)

Copied to clipboard

Challenge: Existing attempts to model the relationship between the real world and written or spoken text have focused on more interpretable and simplistic text representations.
Approach: They propose to link shifts in semantic embedding space to real-world market shocks and partisanship to shape predictions of market fluctuations.
Outcome: The proposed model demonstrates that partisanship can influence the predictive power of text for market fluctuations and shape reactions to those same shocks.
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks (2025.acl-long)

Copied to clipboard

Challenge: a study aims to assess the fairness and robustness of Large Language Models in dialectal queries . speakers of "non-standard" dialects are known to experience implicit and explicit discrimination .
Approach: They propose to use a benchmark to assess the fairness of large language models in dialects . they hire speakers with computer science backgrounds to rewrite seven popular benchmarks based on AAVE .
Outcome: The proposed benchmarks show that most models show significant brittleness and unfairness to queries in AAVE.
ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts (2025.emnlp-main)

Copied to clipboard

Challenge: Scientific fact-checking has largely focused on textual and tabular sources, neglecting scientific charts.
Approach: They propose a benchmark for scientific fact-checking grounded in scientific charts . climateViz comprises 49,862 claims paired with 2,896 visualizations . results show current models struggle to perform fact- checking when statistical reasoning is required .
Outcome: The climateviz benchmark is the first large-scale benchmark for scientific fact-checking . it includes 49,862 claims paired with 2,896 visualizations labeled as support, refute, or not enough .
Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics (2024.lrec-main)

Copied to clipboard

Challenge: Scalar adjectives describe different domain scales and vary in intensity . they can be triggered by scalar adjective and require listeners to reason pragmatically about them.
Approach: They probe different families of Large Language Models for their knowledge of the lexical semantics of scalar adjectives and one specific aspect of their pragmatics.
Outcome: The proposed models encode rich lexical-semantic information about scalar adjectives but lack a good understanding of skalar diversity.
Quantifying Compositionality of Classic and State-of-the-Art Embeddings (2025.findings-emnlp)

Copied to clipboard

Challenge: Static word embeddings make strong claims about compositionality, but the SOTA generative models go too far in the other direction.
Approach: a new study evaluates the compositionality of word embeddings by canonical correlation analysis . strong compositional signals are observed in later training stages across data modalities .
Outcome: a new evaluation of compositional models shows that they exploit access meanings when justified . strong compositional signals are observed in later training stages and in deeper layers of the transformer-based model before a decline at the top layer.
Actors, Frames and Arguments: A Multi-Decade Computational Analysis of Climate Discourse in Financial News using Large Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: a new study examines how financial news media portrays climate change . financial news is the nervous system of the global economy .
Approach: They propose a three-stage Actor–Frame–Argument pipeline that uses large language models to extract actors, stances, frames, and argumentative structures from a 980,061-article corpus.
Outcome: The proposed pipeline extracts actors, stances, frames, and argumentative structures from a 980,061-article corpus of climate-related financial news from the Dow Jones Newswire (2000–2023) it is based on a human-annotated gold standard and a Decompositional Verification Framework (DVF) that decomposes evaluation into completeness, faithfulness, coherence, and relevance, with multi-judge scoring calibrated against human ratings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations